Papers with large-scale evaluation dataset
TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route (2025.emnlp-main)
Copied to clipboard
Hongyi Luo, Qing Cheng, Daniel Matos, Hari Krishna Gadi, Yanfeng Zhang, Lu Liu, Yongliang Wang, Niclas Zeller, Daniel Cremers, Liqiu Meng
| Challenge: | Existing studies on large language models have limited evaluation of their geospatial cognition . a unified framework for evaluating geospcial cognition in LLMs remains absent . |
| Approach: | They propose a benchmark to evaluate the geospatial route cognition of Large Language Models . they propose 'pathbuilder' tool for converting natural language instructions into navigation routes . |
| Outcome: | The proposed framework and metrics evaluate 9 state-of-the-art LLMs on route reversal task. |
PERSONA: A Reproducible Testbed for Pluralistic Alignment (2025.coling-main)
Copied to clipboard
| Challenge: | Currently, preference optimization approaches fail to capture the plurality of user opinions . Currently used methods do not account for the pluralities of users and difference of opinion . |
| Approach: | They propose a reproducible test bed to evaluate pluralistic alignment of language models . they generate user profiles from census data and use a large-scale evaluation dataset . |
| Outcome: | The proposed model improves pluralistic alignment of language models with diverse user values . it generates a large-scale evaluation dataset with 317,200 feedback pairs . |